Skip to content

[WS1][CUDA][MiniMax-H3] ws1_one_h3_block: one real block forward/backward with first-drift report (RFC #420 stack 10/10) - #539

Open
fusheng-ji wants to merge 148 commits into
RL-Align:test-h3from
fusheng-ji:feat/h3-one-block
Open

fusheng-ji wants to merge 148 commits into
RL-Align:test-h3from
fusheng-ji:feat/h3-one-block

Conversation

@fusheng-ji

Copy link
Copy Markdown

Draft. Opened before the claim on #420; it will be marked ready once the row is assigned. Stack: #483 → … → #494 → this PR, so it merges last. This PR's own diff has two commits: the code commit and the evidence commit.

Status

No known failures at 41e4293 (code) / caaf754 (evidence). 11 of the 17 block nodes run on interim operators or the provider replay until their rows land (see Bindings). The block-level promises below hold with those bindings.

What

RFC #420 row ws1_one_h3_block: one real H3 block, forward and backward, with every intermediate and a first-drift report.

  • The block. Block 0 of MiniMax-H3@42ed227 runs as a 17-node graph, from the AdaLN projection to the second gated residual. It runs on packed FL2VA layouts whose position_ids and tags are bitwise equal to the pinned diffusers pipeline's build_packed_sequence; the hashes are pinned in a CPU test.
  • Three executions per node. Each node runs as:
    • the diffusers replay (op for op, no diffusers import);
    • the RL-Kernel binding, both chained and isolated (fed diffusers' inputs);
    • an FP32 golden.
  • Ownership. Each node names the RFC row that owns its arithmetic, and whether it is bound to that row's operator, an interim deterministic operator, or the provider replay.
  • Forward report, per node:
    • repeat and batch-row bitwise checks;
    • chained and isolated agreement with diffusers (including the first differing packed row);
    • candidate and diffusers error against the golden;
    • the node's sha256;
    • first_{,isolated_,repeat_,batch_}drift.
  • Backward report:
    • gradients of the block input, temb and all 12 block parameters (repeat, error vs golden);
    • the gradient reaching every node (repeat, batch row), with first_grad_{repeat,batch}_drift.
  • Weights. Block-0 attention/FFN weights are pinned in a new manifest section (block_tensors). They are extracted by prepare_h3_weights.py --block from the same shard, so the existing conditioning file is unchanged.

Bindings

Node Owning row Bound to
adaln_projection adaln_projection_3mod row (#485), via the registry
norm1, norm2 h3_rmsnorm row (#487), via the registry
residual_attn, residual_mlp adaln_gate_residual row (#488), via the registry
q/k/v_proj, o_proj, ffn_gate_up, ffn_down h3_qkv_gemm, h3_attention_o_gemm, h3_ffn_gate_up_gemm, h3_ffn_down_gemm interim: Triton FP32-accumulate GEMM, one BF16 rounding
q_norm, k_norm h3_qk_rmsnorm_d128 interim: #487's plain H3RMSNormCudaOp on the 128-wide heads (bitwise equal to nn.RMSNorm)
rope_q, rope_k h3_mm_rope_3axis_partial reference: the provider replay, since no MM-RoPE kernel exists yet
attention h3_full_attention interim: DeterministicAttentionOp(causal=False)
swiglu h3_swiglu interim: SwiGLUCudaOp

When an owning row lands, swapping its node's candidate callable is the whole integration, and the report entries become that operator's block-level evidence.

Prior art & reuse decision

This is a validation closeout, not a new kernel; no operator is written here.

  • diffusers. MiniMaxH3TransformerBlock is the executable reference. It is replayed op for op as the provider (h3_provider), which pins its dtypes and op order without importing diffusers.
  • [ws1]: WS1 Full Qwen3-8B Dense Train-Inference Closeout #315 Qwen3-8B chain gate (rl_engine/validation/models/chain_gate.py). Its method is reused: per-node digests and first drift in graph order. Its graph is Qwen3's (causal, no AdaLN, Qwen shapes), so the H3 block is built on the h3_chain stage pattern of [WS1][CUDA][MiniMax-H3] timestep_sinusoid_h3: FP32 timestep features + shared H3 harness (RFC #420 stack 1/9) #483.
  • Interim operators are reused from RL-Kernel. The one decision taken here is the GEMM entry point. TritonDetGemmOp.__call__ reduces K as a tree with BF16 nodes; its FP32-accumulate kernel is equally deterministic and row-invariant, matches cuBLAS's accuracy, and is 9–17× faster (B200, error vs FP64):
Shape Tree path FP32 K loop → BF16 cuBLAS Tree time FP32 K loop time
264 × 5376 → 7168 7.0e-3 2.4e-3 2.4e-3 10.4 ms 1.2 ms
925 × 14336 → 5376 7.5e-3 2.6e-3 2.6e-3 47 ms 2.7 ms

Results

Generated from a clean tree at 41e4293 on an otherwise idle B200 (torch 2.13.0+cu130), with tools/validation/models/h3_block_replay.py --layouts tiny,small,medium: report.json.

one-block evidence

S First drift (chained / isolated) Repeat or batch drift, any node or gradient Output err: RL-Kernel / diffusers Fwd ms Fwd+bwd ms
264 adaln_projection / adaln_projection none 5.21e-3 / 5.21e-3 7.4 vs 0.61 22.7 vs 2.8
925 adaln_projection / adaln_projection none 7.87e-3 / 7.87e-3 30.1 vs 1.12 81.7 vs 5.0
3160 adaln_projection / adaln_projection none 8.97e-3 / 8.97e-3 225 vs 3.8 577 vs 11.5
  • Isolated bitwise promises. Every node that promises it is bitwise equal to diffusers on diffusers' inputs at every size: norm1/norm2, q_norm/k_norm, RoPE, and both gated residuals. The first drift is the first reduction whose tree differs from cuBLAS.
  • Forward accuracy. Per-node error vs the FP32 golden is 0.93–1.20× diffusers'. The largest is the interim attention at S = 3160.
  • Gradient accuracy:
    • The candidate is up to 7× more accurate than diffusers on adaln_proj.weight, adaln_proj.bias and temb. diffusers' index_select backward accumulates in BF16 atomics; the row operators use FP32 segment sums.
    • The candidate is within 1.16× on the other leaves, except attn.norm_q.weight at S = 3160 (9.5e-3 vs 5.5e-3). At S = 925 the same leaf is 2× more accurate than diffusers, so the 128-element gradient's error follows BF16 rounding, not one side.
  • Control H-C3, reported, not promised. Permuting the packed rows with their positions and indices changes 79 / 450 / 2199 output rows (max 8). Attention sums keys in packed order; settling this belongs to h3_full_attention.
  • Speed. The 12–60× forward cost is the interim GEMMs and attention (FP32 scores materialised), not the row operators.

Tests

RL_KERNEL_H3_WEIGHTS=<dir> CUDA_VISIBLE_DEVICES=0,1 python -m pytest tests/models/minimax_h3 -q

501 passed, 5 skipped on two B200s, including the 9 new tests in test_h3_one_block.py. The new tests check:

  • CPU:
    • the layouts are pinned to diffusers;
    • the graph is well formed;
    • MM-RoPE rotates 96 of 128 channels.
  • GPU forward (264, 925):
    • every node repeats and is batch-row bitwise;
    • promised nodes are isolated-bitwise;
    • node and output errors are at most 1.25 × diffusers' + 1e-3.
  • GPU backward (264):
    • every leaf repeats;
    • dX is batch-row bitwise;
    • there is no gradient drift;
    • leaf errors are at most 1.5 × diffusers' + 5e-3.

Scope / Known limitations

…+ shared H3 harness

RFC RL-Align#420 step 3, row `timestep_sinusoid_h3`.

- CUDA kernel: one thread per (t, k), the provider's exact FP32 op
  sequence (__fmul_rn, precise expf/sinf/cosf); bitwise equal to
  diffusers' get_timestep_embedding on CUDA. Analytic row-local backward.
  Non-1-D, empty, integer, non-finite or out-of-[0, 1] timesteps fail
  closed (RFC probe H10).
- PyTorch reference (provider replay + FP64 golden), registry, gtest.
- Shared H3 harness used by the following rows:
  * h3_manifest.json + prepare_h3_weights.py: pinned revision and sha256,
    only shard 1 of 14 is downloaded;
  * h3_provider.py: diffusers replayed op for op;
  * h3_chain.py: stage-wise chain replay, each stage declaring its
    promises (provider-bitwise, golden tolerance); used by
    scripts/h3_chain_replay.py and tests/h3/test_h3_conditioning_e2e.py;
  * h3_report.py + scripts/h3_evidence.py + scripts/plot_h3_evidence.py:
    per-op performance/accuracy JSON and figure;
  * benchmarks/benchmark_h3_conditioning.py.

Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Produced by scripts/h3_evidence.py at 0522865 on a clean tree; rendered with scripts/plot_h3_evidence.py and embedded in the operator doc.

Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
…p MLP

RFC RL-Align#420 step 3, row `timestep_mlp_fp32` (stacked on timestep_sinusoid_h3).

- New CUDA module csrc/cuda/h3/det_linear.cu, contract h3-det-linear-v1:
  warp per output-column group, lane-strided 16-byte K chunks, ascending
  fmaf chains, xor butterfly, bias then fused SiLU. The launch shape never
  changes a column's summation order, so rows are batch/position
  invariant and runs are repeat-bitwise.
- Backward without cuBLAS or atomics: dinput via fixed 64-row N chunks
  folded in order (row-local), dweight/dbias as ascending-row FP32 folds.
- PyTorch reference (provider path + FP64 golden); BF16 anywhere is
  rejected (RFC probe H7: the time_embedder is a declared FP32 module).
- On the pinned weights both CUDA and cuBLAS are ~1e-6 from the FP64
  golden (100x inside the contract); the op is 2x faster for T >= 2.
- Chain stage (covered by the end-to-end test), report/figure hooks,
  registry, gtest, doc.

Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Produced by scripts/h3_evidence.py at 65ef7f6 on a clean tree; rendered with scripts/plot_h3_evidence.py and embedded in the operator doc.

Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
…rojection

RFC RL-Align#420 step 3, row `adaln_projection_3mod` (stacked on timestep_mlp_fp32).

- SiLU at temb's FP32 precision and exactly one cast to the projection
  dtype (bitwise equal to the provider activation); a BF16 temb is
  rejected (RFC probe H7).
- 2688 -> 96768 projection on a BF16 tensor-core path in det_linear.cu,
  contract h3-det-linear-bf16-mma-v1: warp = 16 output columns x 8 input
  rows (zero-padded), K in ascending 16-wide groups, each mma.sync
  m16n8k16 from a zero accumulator added into an FP32 running sum. Same
  instruction sequence for every T <= 8, so rows are bitwise
  batch/order/tile invariant; 99.98% of outputs correctly rounded vs
  99.84% for cuBLAS; kernel time flat at ~78 us for T = 1..8. SM80+.
- Six (3T, 5376) views in diffusers' view(-1, 6H).chunk(6) layout, row
  t*3+m; a channel-tagging test pins chunk/modality placement (probe H4).
- Backward: FP32 d_temb through the identity-VJP cast, ascending-row
  dW/db folds, no cuBLAS or atomics. The FP64 golden applies the cast
  straight-through so its own gradient is not rounded to BF16.
- Chain stage (end-to-end test), report/figure hooks, registry, gtest, doc.

Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Produced by scripts/h3_evidence.py at d06e120 on a clean tree; rendered with scripts/plot_h3_evidence.py and embedded in the operator doc.

Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
RFC RL-Align#420 step 3, row `adaln_row_gather` (stacked on adaln_projection_3mod).

- Forward: one launch gathers all six (S, H) modulation tensors by
  timestep_index * 3 + token_tag with 16-byte copies; bitwise equal to
  diffusers' six index_select calls. Out-of-range tags or timesteps fail
  closed (RFC probes H2/H3), never clamped.
- Backward: deterministic FP32 segmented sum (stable sort by row, fixed
  256-position tiles, ordered fold, one cast; no atomics) instead of
  index_select's atomic BF16 scatter-add. The PyTorch reference gets an
  equally deterministic backward.
- H3AdaLNModulationCudaOp: projection + gather as one autograd node,
  bitwise-equal forward, FP32 table gradient into the projection backward.
- Chain replay gains the gather stage and a backward mode (separate,
  fused and diffusers chains vs an FP64 golden); the end-to-end test now
  covers the whole chain forward and backward.
- Report/figure hooks, registry, gtest (reduction class: the VJP is a
  segmented sum), doc.

Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
…y and figure

Produced by scripts/h3_evidence.py and scripts/h3_chain_replay.py --backward at fa551c2 on a clean tree; rendered with scripts/plot_h3_evidence.py and embedded in the operator doc.

Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
RFC RL-Align#420 step 4, row `h3_rmsnorm` (stacked on adaln_row_gather).

- CUDA kernel replaying PyTorch's own RMSNorm reduction order
  (vectorized_layer_norm_kernel: (32, 4) block per row, 4-element
  vectors, shuffle-down and cross-warp trees, rsqrtf, w * (rstd * x)):
  bitwise equal to nn.RMSNorm on all four pinned H3 norm weights and
  batch/position invariant.
- Fused AdaLN modulation n * (1 + scale[i]) + shift[i] with the rows
  gathered in-kernel from strided table views and each step rounded where
  the eager expression rounds: bitwise equal to diffusers for the block
  (adaln_indices) and norm_out (timestep_indices), without materialising
  the (S, H) gathered tensors; 2.5x faster forward at S = 131072.
- Deterministic backward: row-local dx, dweight over fixed 256-row tiles,
  dshift/dscale as sorted segment sums (diffusers: non-deterministic BF16
  atomics, ~15x larger error).
- Manifest pins the block/refiner-final/norm_out norm weights and
  norm_out.linear (all in shard 1); chain gains a per-mode history and the
  block norm1 + modulation stage (covered by the end-to-end test).
- Registry, gtest, report/figure hooks, doc.

Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Produced by scripts/h3_evidence.py at f36aa60 on a clean tree; rendered with scripts/plot_h3_evidence.py and embedded in the operator doc.

Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
…ernel gate gather

RFC RL-Align#420 step 6, row `adaln_gate_residual` (stacked on h3_rmsnorm).

- residual + gate[adaln_index] * sublayer_output in one pass, the gate row
  read from the AdaLN table view in the kernel; the two roundings sit where
  the eager expression rounds (__fmul_rn/__fadd_rn, no FMA contraction), so
  the forward is bitwise equal to diffusers in bf16/fp16/fp32. 1.7x faster
  forward at S = 131072.
- Backward: d_residual = grad; d_sublayer bitwise equal to the eager VJP
  (exact product, one rounding); d_gate a deterministic FP32 segment sum
  (diffusers: non-deterministic BF16 atomics, 13-49x larger error, varying
  run to run). The PyTorch reference gets an equally deterministic gate
  gradient.
- Chain stage residual + gate_msa * (stand-in attention output), covered
  by the end-to-end test; registry, gtest, report/figure hooks, doc.

Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Produced by scripts/h3_evidence.py at a26b41a on a clean tree; rendered with scripts/plot_h3_evidence.py and embedded in the operator doc.

Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
…norm/modulation

RFC RL-Align#420 step 7, row `final_adaln_out` (stacked on adaln_gate_residual).

- norm_out as one autograd node: FP32 SiLU -> one BF16 cast -> the
  deterministic tensor-core GEMV on norm_out.linear (shift first), then
  the h3_rmsnorm kernel indexed by timestep_indices. Forward equal to
  diffusers on 99.96% of elements (1-ULP projection ties); rows bitwise
  batch/position invariant; 2.4x faster forward at S = 131072.
- Backward keeps the table gradient in FP32 into the projection backward:
  d_temb/dW/db 13-57x closer to FP64 than diffusers (whose error varies
  run to run), all gradients repeat-bitwise.
- Golden rounds only at the declared boundaries (SiLU cast, BF16 table),
  straight-through for the gradient; gtest inputs at realistic scale.
- Chain stage: the gated residual stream -> norm_out with the MLP stage's
  temb (end-to-end test now runs timestep -> ... -> norm_out); registry,
  report/figure hooks, doc.

Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
… and figure

Produced by scripts/h3_evidence.py and scripts/h3_chain_replay.py --backward at 38d575f on a clean tree; rendered with scripts/plot_h3_evidence.py and embedded in the operator doc.

Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>

# Conflicts:
#	csrc/ops.cpp
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>

# Conflicts:
#	scripts/plot_h3_evidence.py
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Validate pinned tensor contents at preparation and load time. Execute chain prerequisites for selected stages and require checkpoint weights for weighted evidence. Reject invalid CUDA tile and backward inputs, align the CUDA wrapper guards, skip unsupported devices, and rotate comparative timing with raw samples.

Validation: 143 H3 tests passed on B200 after a fresh CUDA build; illegal FP32 tiles fail compilation; isolated AdaLN replay and pinned-weight evidence generation passed; Black, isort, Ruff, and Flake8 passed.
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Check timestep and modality bounds on the device before computing gather addresses, including unchecked calls and overflowing int64 indices. Reject non-prefix replay stage selections and partial backward chains before loading weights.

Validation: rebuilt the CUDA extension on B200; tests/h3: 175 passed, including 12 isolated native bounds probes and 16 CLI regressions.
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
CodeRabbit skips automatic reviews targeting non-default branches by default. Allow the test-h3 base branch so pushes to the H3 stack receive incremental reviews.

Validation: configuration validated against the official CodeRabbit JSON schema.
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Include the RFC 420 target branch in CodeRabbit automatic reviews so pushes to the H3 stack are reviewed instead of silently skipped.

Validation: configuration parses and validates against the published CodeRabbit v2 schema; the additional branch pattern matches only test-h3.
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
RFC RL-Align#420 row ws1_one_h3_block. Block 0 of the pinned checkpoint runs as a
17-node graph (AdaLN projection through the second gated residual) on
packed FL2VA layouts that are bitwise equal to the pinned diffusers
pipeline's. Each node runs as the diffusers replay, the RL-Kernel binding
(chained and on provider inputs) and an FP32 golden, and names the RFC row
that owns it and whether it is bound to that row's operator, an interim
deterministic operator, or the provider replay.

The report gives per node: repeat and batch-row bitwise checks, provider
agreement, error against the golden for candidate and provider, sha256,
and the first drift of each comparison; the backward replay gives leaf
gradients (block input, temb, 12 parameters) and the first node whose
incoming gradient drifts. A permutation probe reports control H-C3.

The interim GEMM nodes use the Triton FP32-accumulate kernel with one BF16
rounding: TritonDetGemmOp's default tree path rounds every 32-wide leaf to
BF16 and has about 3x cuBLAS's error at these shapes.

Block-0 attention/FFN weights are pinned in a new manifest section and
extracted with prepare_h3_weights.py --block.

Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
Forward and backward replay of block 0 at S = 264, 925 and 3160, written
from a clean tree at 41e4293: no repeat or batch-row drift at any node or
gradient; first drift from diffusers is adaln_projection; block output as
close to the FP32 golden as diffusers'. Adds the figure script.

Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
@coderabbitai

coderabbitai Bot commented Oct 11, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 19f7c3ed-5943-432e-94e1-026c5d3f2a64

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant